Papers with Vision Language Models
Beyond Visual Understanding Introducing PARROT-360V for Vision Language Model Benchmarking (2025.coling-industry)
Copied to clipboard
| Challenge: | Current benchmarks for evaluating Vision Language Models (VLMs) often fail to thoroughly assess these models’ abilities to understand complex visual and textual content. |
| Approach: | They propose a benchmark that features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks. |
| Outcome: | The PARROT-360V Benchmark features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks. |
CoLLaVO: Crayon Large Language and Vision mOdel (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing Large Language Models (LLMs) and instruction tuning have been used to drive the evolution of Vision Language Model (VLM) towards a versatile general-purpose model. |
| Approach: | They propose a learning strategy of Dual QLoRA to preserve object-level image understanding without forgetting it during visual instruction tuning, thereby achieving a significant leap in numerous VL benchmarks in a zero-shot setting. |
| Outcome: | The proposed model outperforms closed-source models on vision language tasks and achieves a significant leap in numerous benchmarks. |
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine (2025.naacl-long)
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) have demonstrated promise in generating visually grounded responses, but their application in the medical domain is hindered by unique challenges. |
| Approach: | They propose a vision language model with versatile visual grounding for medicine that generates semantic segmentation masks and instance-level bounding boxes. |
| Outcome: | The proposed model can generate semantic segmentation masks and instance-level bounding boxes, and accommodates various imaging modalities, including both 2D and 3D data. |
MATE: Meet At The Embedding - Connecting Images with Long Texts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Vision Language Models (VLMs) focus on aligning images with short descriptive captions. |
| Approach: | They propose a method that combines VLMs with Large Language Models to efficiently align images with long texts without additional text pairs. |
| Outcome: | The proposed method bridges the gap between VLM and LLM without additional image-long text pairs. |
CAST: Cross-modal Alignment Similarity Test for Vision Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) are typically evaluated with Visual Question Answering tasks which assess a model’s understanding of scenes. |
| Approach: | They propose to use visual question answering (VQA) to assess a model's understanding of scenes to probe for self-consistency across modalities. |
| Outcome: | The proposed test does not focus on objective accuracy but rather on whether VLMs are internally consistent in their outputs. |
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning (2025.emnlp-main)
Copied to clipboard
Mingyuan Wu, Jize Jiang, Haozhen Zheng, Meitang Li, Zhaoheng Li, Beitong Tian, Bo Chen, Yongjoo Park, Minjia Zhang, ChengXiang Zhai, Klara Nahrstedt
| Challenge: | Recent Vision Language Models (VLMs) have shown tremendous promise in a wide range of realworld applications, but their size has made at-scale deployment and operation challenging due to high consumption of cloud computing resource, high latency, and expensive API calls. |
| Approach: | They propose a master–apprentice framework for collaborative inference between large and small vision language models. |
| Outcome: | The proposed framework improves reasoning performance on widely-recognized and challenging general reasoning benchmarks and specifically boosts reasoning of apprentice VLMs by 36.6%. |
Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone (2026.findings-eacl)
Copied to clipboard
| Challenge: | Using vision language models, we examine demographic biases in VLMs across gender, race, age, and skin tone. |
| Approach: | They propose a benchmark for uncovering demographic biases in Vision Language Models . they propose 'Gras Bias Score' to quantify bias in VLMs based on gender, race, age and skin tone . |
| Outcome: | The proposed model achieves 98, far from the unbiased ideal of 0. |
Defeating Cerberus: Privacy-Leakage Mitigation in Vision Language Models (2026.findings-eacl)
Copied to clipboard
Boyang Zhang, Istemi Ekin Akkus, Ruichuan Chen, Alice Dethise, Klaus Satzke, Ivica Rimac, Yang Zhang
| Challenge: | Existing models that process multiple modalities of data have been used for multimodal tasks, but their advanced capabilities raise privacy concerns. |
| Approach: | They propose a method to modify the model’s internal states associated with PII-related content and to reduce the risk of PI I leakage by modifying the model's internal state. |
| Outcome: | The proposed method achieves on average 93.3% refusal rate for various PII-related tasks with minimal impact on unrelated model performances. |
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models and Vision Language Model (VLMs) have demonstrated aptitude as potential substitutes for human participants in psycholinguistic experiments. |
| Approach: | They examine whether large language models and vision language models implicitly understand sound-based phenomena via orthography and imagery alone. |
| Outcome: | The proposed models demonstrate sound symbolism and ability to "hear" using language and vision modules. |
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines (2025.naacl-long)
Copied to clipboard
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo
| Challenge: | Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts. |
| Approach: | They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset. |
| Outcome: | The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages. |
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing research on UI/UX automation often requires high-fidelity inputs like Figma designs or detailed screenshots, limiting accessibility and impeding efficient design iteration. |
| Approach: | They propose a benchmark that evaluates state-of-the-art Vision Language Models on converting sketches into webpage prototypes. |
| Outcome: | The benchmark evaluates state-of-the-art Vision Language Models on automating the conversion of rudimentary sketches into webpage prototypes. |
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding (2025.findings-acl)
Copied to clipboard
| Challenge: | Vision Language Models struggle with visual arithmetic, seemingly simple tasks like object counting or length comparison, which are essential for relevant complex tasks like chart understanding and geometric reasoning. |
| Approach: | They propose a novel post-training strategy inspired by Piaget’s theory of cognitive development that trains VLMs to recognize invariant properties under visual transformations. |
| Outcome: | The proposed approach outperforms supervised fine-tuning methods while requiring 60% less training data. |
Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies (2025.coling-main)
Copied to clipboard
| Challenge: | Existing fact-checking systems that use text and image information are susceptible to fake news spread by social media platforms. |
| Approach: | They propose a neural probing classifier based on multimodality and embeddings from text and image encoders to represent multimodal content for fact-checking. |
| Outcome: | The proposed classifier outperforms KNN and SVM baselines in leveraging extracted embeddings, highlighting its effectiveness for multimodal fact-checking. |
Benchmarking Vision Language Models for Cultural Understanding (2024.emnlp-main)
Copied to clipboard
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Steenkiste, Lisa Hendricks, Karolina Stanczak, Aishwarya Agrawal
| Challenge: | Recent multimodal vision-language models have shown impressive performance in tasks such as image-to-text generation, visual question answering, and image captioning. |
| Approach: | They propose a visual question-answering benchmark to assess VLMs' cultural understanding of various facets of culture from 11 countries across 5 continents. |
| Outcome: | The visual question-answering benchmark aims to assess VLMs' cultural understanding across regions. |
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch (2026.acl-long)
Copied to clipboard
Zheng Liu, Honglin Lin, Xiaoyang Wang, Xin Gao, Yu Li, Mengzhang Cai, Yun Zhu, Zhanping Zhong, Qizhi Pei, Zhuoshi Pan, Xiaoran Shang, Conghui He, Bin Cui, Wentao Zhang, Lijun Wu
| Challenge: | Existing open-source vision language models lack high-quality training data for chart reasoning . current models are simplistic and repetitive, while associated QA pairs are prone to hallucinations . |
| Approach: | They propose a framework to synthesize complex charts and reliable reasoning data from scratch. |
| Outcome: | Experimental results show that ChartVerse-8B surpasses existing models in QA and difficulty . lack of high-quality training data hampers development of open-source models . |
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current models struggle with long-form videos due to the quadratic complexity of attention mechanisms. |
| Approach: | They propose a model-agnostic framework that leverages temporal cues from queries to prune video tokens. |
| Outcome: | The proposed framework reduces computation by 65% while preserving 97-99% of original performance. |
PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of PII leakage ignore how a subject’s online presence affects privacy alignment. |
| Approach: | They propose a benchmark that evaluates safety through the continuum of online presence by stratifying 200 subjects into four visibility categories: high, medium, low, and zero. |
| Outcome: | The proposed model stratifies 200 subjects into four visibility categories based on the extent and nature of their information available online. |
Beyond Screenshots: Evaluating VLMs’ Understanding of UI Animations (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent studies of Vision Language Models (VLMs) for UI understanding have focused primarily on static screenshots, leaving it unclear how well these models handle dynamic UI animations. |
| Approach: | They evaluate UI animation models' ability to perceive animation effects and interpret animation meaning . they use motion, context, and perceptual cues to probe factors affecting VLM performance . |
| Outcome: | The proposed model can detect primitive motion, but its interpretation is inconsistent . the proposed model is based on 300 annotated UI animation videos . |
Cultivating Gaming Sense for Yourself: Making VLMs Gaming Experts (2025.acl-long)
Copied to clipboard
| Challenge: | Recent efforts leverage Vision Language Models (VLMs) as direct controllers, often pausing the game to analyze screens and plan action through language reasoning. |
| Approach: | They propose a paradigm shift in gameplay agent design that uses Vision Language Models as a developer instead of direct control. |
| Outcome: | The proposed framework achieves fluent gameplay in diverse genres, including ACT, FPS, and Flappy Bird, setting a new benchmark for game-playing agents. |
Grounding Task Assistance with Multimodal Cues from a Single Demonstration (2025.findings-acl)
Copied to clipboard
Gabriel Herbert Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet, Andrew D Wilson
| Challenge: | RGB video often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior. |
| Approach: | They propose a framework that integrates eye gaze and speech cues to improve conversational agents for task assistance by integrating eye gaze with speech cuests. |
| Outcome: | The proposed framework captures fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering. |
VADE: Visual Attention Guided Hallucination Detection and Elimination (2025.findings-acl)
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) are prone to hallucinations, generating outputs that lack grounding in the actual visual data. |
| Approach: | They propose a sequence modelling approach to learn complex sequential patterns from transformer attention maps. |
| Outcome: | The proposed approach achieves an average PR-AUC of 80% in hallucination detection on M-HalDetect and an 5% improvement in hallucinosis mitigation on MSCOCO. |
Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing work on 3D radiograph report generation focuses on 2D images, but 3D medical images provide more comprehensive diagnostic information. |
| Approach: | They propose a comprehensive training recipe for building high-performing VLMs for 3DRRG using a publicly available 3D CT-report dataset. |
| Outcome: | The proposed model achieves superior performance across different model sizes and input 3D medical image resolutions. |
Iterative Prompt Refinement for Safer Text-to-Image Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing safety methods for text-to-image models ignore the images produced . this can result in unsafe outputs or unnecessary changes to already safe prompts . |
| Approach: | They propose an iterative prompt refinement algorithm that uses Vision Language Models to analyze prompts and generated images. |
| Outcome: | The proposed method improves safety while maintaining user intent and reliability comparable to existing methods. |
Granular Privacy Control for Geolocation with Vision Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions. |
| Approach: | They develop a benchmark to evaluate the ability of VLMs to moderate geolocation dialogues with users. |
| Outcome: | a new benchmark evaluates the ability of VLMs to moderate geolocation conversations with users. |
GUICourse: From General Vision Language Model to Versatile GUI Agent (2025.acl-long)
Copied to clipboard
Wentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin, Junjie Fang, Yue Zhao, Chongyi Wang, Jun Liu, Guirong Chen, Yupeng Huo, Yuan Yao, Yankai Lin, Zhiyuan Liu, Maosong Sun
| Challenge: | Graphical User Interfaces (GUIs) are a pivotal medium for human-computer interaction. |
| Approach: | They propose a series of datasets for training visual-based GUI agents using general VLMs. |
| Outcome: | The proposed GUICourse datasets show that even a small-sized GUI agent performs better on GUI tasks. |
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation (2025.findings-emnlp)
Copied to clipboard
Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Vladimir Araujo, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofía Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio
| Challenge: | a human-curated benchmark of over 5,800 triples of images is used to evaluate multimodal translation systems. |
| Approach: | They introduce a human-curated benchmark of over 5,800 triples of images along with parallel captions in English and regional languages. |
| Outcome: | The results show that visual context improves translation quality in culturally-specific items . |
VIBE: Can a VLM Read the Room? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Vision Language Models (LLMs) cannot account for the role that non-verbal cues play in understanding social situations. |
| Approach: | They propose a task to test the capabilities of Vision Language Models (VLMs) to account for the visual social-pragmatic inference gap. |
| Outcome: | The proposed task tests the capabilities of a VLM for a social reasoning task. |
RG-VQA: Leveraging Retriever-Generator Pipelines for Knowledge Intensive Visual Question Answering (2025.findings-emnlp)
Copied to clipboard
Settaluri Lakshmi Sravanthi, Pulkit Agarwal, Debjyoti Mondal, Rituraj Singh, Subhadarshi Panda, Ankit Mishra, Kiran Pradeep, Srihari K B, Godawari Sudhakar Rao, Pushpak Bhattacharyya
| Challenge: | Existing methods to improve the reasoning capabilities of VQA systems are limited due to complexity of graph neural networks and end-to-end training. |
| Approach: | They propose a method to integrate Dense Passage Retrievers with Vision Language Models to boost the reasoning capabilities of VQA systems. |
| Outcome: | The proposed method outperforms human accuracy and GPT-4 in the ScienceQA dataset. |
Unlocking Human-Like Visible Logic: How Logic Diagrams Boost Logic Reasoning in Large Language Models? (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have demonstrated their remarkable capabilities in natural language understanding and generation, but they struggle with formal logical reasoning. |
| Approach: | They propose to incorporate visual logic diagrams into LLMs’ reasoning workflows to enhance their performance on formal logic tasks. |
| Outcome: | The proposed model improves on syllogistic and conditional reasoning with programmatically generated Venn, Euler, and Linear diagrams. |
VISaGE: Understanding Visual Generics and Exceptions (2025.emnlp-main)
Copied to clipboard
| Challenge: | atypical evaluation instances disrupt incontext instance understanding and in-weight conceptual knowledge. |
| Approach: | They propose to use a dataset to analyze atypical visual and textual images to test their models. |
| Outcome: | The proposed model is based on a dataset consisting of typical and exceptional images. |
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches for grounding events in videos are limited by their time-sensitive nature . arrow of time in physics characterizes intrinsic directionality of temporal processes . |
| Approach: | They propose a framework that explicitly models temporal directionality in events to improve event grounding and temporal understanding in VLMs. |
| Outcome: | The proposed framework improves event grounding and directionality understanding in VLMs. |
Bears, all bears, and some bears. Language Constraints on Language Models’ Inductive Inferences (2026.findings-acl)
Copied to clipboard
| Challenge: | Language places subtle constraints on how we make inductive inferences. |
| Approach: | They propose to use language to constrain inductive inferences by replicating an experiment . they find subtle differences arise in general purpose statistical learners like VLMs . |
| Outcome: | The proposed model can be used to extend inductive inferences to humans using language . the model can extend properties of a category to other members of the population, the authors show . |
GeoRC: A Benchmark for Geolocation Reasoning Chains (2026.acl-long)
Copied to clipboard
Mohit Talreja, Joshua Diao, Jim James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays
| Challenge: | Vision Language Models (VLMs) are good at recognizing the global location of a photograph but are startlingly bad at explaining which image evidence led to their location prediction. |
| Approach: | They propose a benchmark for geolocation reasoning chains based on the global location prediction task in the popular GeoGuessr game. |
| Outcome: | The proposed benchmark compares LLM-as-a-judge and VLM-As-jumble strategies against human scoring. |
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Recent work on Chain-of-Thought prompting imposes substantial computational overhead . lack of supervision obscures the analyzability of the latent reasoning chain. |
| Approach: | They propose a framework to render latent reasoning chain into images, making latent rationale explicit and traceable. |
| Outcome: | The proposed framework achieves 3-4 token compression and substantial inference acceleration compared to explicit CoT prompting. |